Day 10 我們建立了第一組 Eval Dataset。
目前 evals/cases.json 裡已經有 15 筆測試案例,包含:
但到目前為止,這些案例還只是靜態資料。
如果要測 Agent,我們還是得手動複製每一題,貼到 python3 app.py 裡執行。
這樣不適合做 Evaluation。
所以 Day 11 開始讓測試流程自動化:
讀取
evals/cases.json,逐筆呼叫 Agent,儲存每題的 Agent output,並綁定對應的 trace session id。
今天還不會判斷 pass / fail。
今天先讓整批測試案例跑起來。
會完成:
今天先不做:
這些從 Day 12 開始逐步處理。
Batch Evaluation Runner 可以想成一個批次測試器。
它負責把 eval dataset 裡的測試案例一題一題丟給 Agent。
流程如下:
讀取 cases.json
-> for each test case
-> 呼叫 agent.run(input)
-> 取得 AgentResult
-> 保存 trace
-> 記錄 output 和 session_id
-> 輸出 evaluation run 結果
今天的 runner 還不會判斷答案是否正確。
所以它不會輸出:
case_001: pass
case_002: fail
而是先輸出:
case_001: completed, session_id=...
case_002: completed, session_id=...
這樣拆,是因為 Evaluation 可以分成兩層。
第一層是執行:
把所有 test cases 跑完,收集 Agent output。
第二層是評分:
根據 expected 和 grading_method 判斷 pass / fail。
Day 11 先處理第一層,Day 12 再處理第二層。
今天新增一個 runner 檔案,並讓 batch run 結果輸出到 data/eval_runs/。
agent-testing-platform/
app.py
trace_viewer_app.py
agents/
__init__.py
simple_agent.py
prompts.py
fake_llm.py
tools/
__init__.py
calculator.py
tracing/
__init__.py
models.py
storage/
__init__.py
database.py
schema.sql
ui/
__init__.py
trace_viewer.py
evals/
__init__.py
cases.json
runner.py
data/
agent_traces.db
eval_runs/
eval_run_*.json
新增:
| 檔案 | 用途 |
|---|---|
evals/runner.py |
讀取 eval dataset,逐筆執行 Agent,輸出 batch run 結果 |
今天不修改:
agents/simple_agent.py
storage/database.py
evals/cases.json
因為 Day 11 的重點是新增 batch runner,不是修改 Agent 行為或測試資料。
執行完一整批 test cases 後,我們會產生一份 JSON 結果。
格式大概長這樣:
{
"run_id": "eval_run_20260906_103000",
"total_cases": 15,
"results": [
{
"case_id": "case_001",
"input": "請計算 135 * 28",
"expected": "3780",
"grading_method": "contains",
"task_type": "calculation",
"status": "completed",
"actual": "The result is 3780",
"trace_session_id": "..."
}
]
}
今天先記錄這些欄位:
| 欄位 | 說明 |
|---|---|
run_id |
這次 batch run 的編號 |
total_cases |
總共跑了幾筆案例 |
case_id |
對應原本的 test case |
input |
Agent 收到的任務 |
expected |
預期答案,今天先保存但不評分 |
grading_method |
評分方式,今天先保存但不使用 |
task_type |
任務類型 |
status |
目前只分成 completed 或 error |
actual |
Agent 實際輸出 |
trace_session_id |
對應到 Trace Viewer 的 session id |
error |
如果執行失敗,記錄錯誤訊息 |
trace_session_id 很有用。
後面看到某一題失敗時,可以用這個 id 回 Trace Viewer 查看當時 Agent 中間做了什麼。
新增 evals/runner.py:
import json
from datetime import datetime
from pathlib import Path
from typing import Any
from agents.fake_llm import FakeLLMClient
from agents.simple_agent import SimpleAgent
from storage.database import init_db, save_trace
BASE_DIR = Path(__file__).resolve().parent.parent
CASES_PATH = BASE_DIR / "evals" / "cases.json"
EVAL_RUNS_DIR = BASE_DIR / "data" / "eval_runs"
def load_cases() -> list[dict[str, Any]]:
with CASES_PATH.open(encoding="utf-8") as file:
return json.load(file)
def create_run_id() -> str:
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
return f"eval_run_{timestamp}"
def run_evaluation() -> dict[str, Any]:
init_db()
agent = SimpleAgent(llm_client=FakeLLMClient())
cases = load_cases()
run_id = create_run_id()
results = []
for test_case in cases:
try:
result = agent.run(test_case["input"])
save_trace(result.trace)
results.append(
{
"case_id": test_case["id"],
"input": test_case["input"],
"expected": test_case["expected"],
"grading_method": test_case["grading_method"],
"task_type": test_case["task_type"],
"status": "completed",
"actual": result.answer,
"trace_session_id": result.trace.session_id,
"error": None,
}
)
except Exception as exc:
results.append(
{
"case_id": test_case["id"],
"input": test_case["input"],
"expected": test_case["expected"],
"grading_method": test_case["grading_method"],
"task_type": test_case["task_type"],
"status": "error",
"actual": None,
"trace_session_id": None,
"error": str(exc),
}
)
eval_run = {
"run_id": run_id,
"created_at": datetime.now().isoformat(),
"total_cases": len(cases),
"results": results,
}
save_eval_run(eval_run)
return eval_run
def save_eval_run(eval_run: dict[str, Any]) -> Path:
EVAL_RUNS_DIR.mkdir(parents=True, exist_ok=True)
output_path = EVAL_RUNS_DIR / f"{eval_run['run_id']}.json"
with output_path.open("w", encoding="utf-8") as file:
json.dump(eval_run, file, ensure_ascii=False, indent=2)
return output_path
def print_summary(eval_run: dict[str, Any]) -> None:
print(f"Run ID: {eval_run['run_id']}")
print(f"Total cases: {eval_run['total_cases']}")
print()
for result in eval_run["results"]:
print(
f"{result['case_id']} | "
f"{result['task_type']} | "
f"{result['status']} | "
f"trace={result['trace_session_id']}"
)
if __name__ == "__main__":
eval_run = run_evaluation()
print_summary(eval_run)
這支檔案就是今天的 Batch Evaluation Runner。
它主要分成幾個 function。
load_cases() 會讀取 evals/cases.json:
def load_cases() -> list[dict[str, Any]]:
with CASES_PATH.open(encoding="utf-8") as file:
return json.load(file)
這裡使用 encoding="utf-8",是因為 cases.json 裡有中文任務。
如果沒有指定編碼,在某些環境中可能會遇到中文讀取問題。
run_evaluation() 是今天最主要的 function。
它會先初始化資料庫:
init_db()
接著建立 Agent:
agent = SimpleAgent(llm_client=FakeLLMClient())
然後讀取所有 cases:
cases = load_cases()
接著逐筆執行:
for test_case in cases:
result = agent.run(test_case["input"])
save_trace(result.trace)
這裡有一個重要設計:
每一筆 test case 執行後,都會把 trace 存進 SQLite。
所以 batch run 的結果會有 trace_session_id,可以對應回 Trace Viewer。
假設 Day 12 之後開始評分,我們看到:
case_013 failed
如果沒有 trace,我們只知道它失敗。
但如果 evaluation result 裡有:
{
"case_id": "case_013",
"trace_session_id": "34ddf4d0-19c3-45d1-b44f-2d5fd6d9b42a"
}
就可以去 Trace Viewer 選這個 session,查看:
這就是 Trace 和 Eval 串起來的地方。
Eval 告訴我們哪一題有問題。
Trace 幫助我們看出問題發生在哪一步。
在 run_evaluation() 裡,我們用 try / except 包住每一筆 test case:
try:
result = agent.run(test_case["input"])
save_trace(result.trace)
...
except Exception as exc:
results.append(
{
"status": "error",
"actual": None,
"trace_session_id": None,
"error": str(exc),
}
)
這樣做是為了避免某一題失敗時,整個 batch run 中斷。
例如某一題讓 Agent 丟出 exception,runner 還是可以繼續跑下一題。
今天只先記錄:
status: error
error: 錯誤訊息
至於這個錯誤是 format_error、tool_error 還是 instruction_error,會留到第三週 Failure Analysis 再處理。
save_eval_run() 會把整次 batch run 結果輸出成 JSON:
def save_eval_run(eval_run: dict[str, Any]) -> Path:
EVAL_RUNS_DIR.mkdir(parents=True, exist_ok=True)
output_path = EVAL_RUNS_DIR / f"{eval_run['run_id']}.json"
with output_path.open("w", encoding="utf-8") as file:
json.dump(eval_run, file, ensure_ascii=False, indent=2)
return output_path
輸出位置會在:
data/eval_runs/
檔名會類似:
eval_run_20260906_103000.json
這樣每次 batch run 都會留下紀錄。
之後 Day 12 加上 evaluator 後,就可以在同一份結果裡加入 pass / fail。
今天還沒有做 dashboard,所以先用終端機印出簡單結果表。
print_summary() 會輸出:
def print_summary(eval_run: dict[str, Any]) -> None:
print(f"Run ID: {eval_run['run_id']}")
print(f"Total cases: {eval_run['total_cases']}")
print()
for result in eval_run["results"]:
print(
f"{result['case_id']} | "
f"{result['task_type']} | "
f"{result['status']} | "
f"trace={result['trace_session_id']}"
)
目前它只顯示:
等 Day 12 開始自動評分後,這裡會再加入:
在專案根目錄執行:
python3 -m evals.runner
預期會看到類似結果:
Run ID: eval_run_20260906_103000
Total cases: 15
case_001 | calculation | completed | trace=8a2b2f21-d1a4-4d8d-9a57-0db43a4e9d23
case_002 | calculation | completed | trace=9c75f807-3c8e-47b2-a38c-55f7b9b2161a
case_003 | calculation | completed | trace=0ad5ef35-334f-44f4-906e-2e95d534f6ac
...
case_015 | json_output | completed | trace=7d734690-0e49-4e3e-8f1e-1d98dbeb820d
如果某一題執行時發生錯誤,會看到:
case_xxx | calculation | error | trace=None
這表示 runner 沒有因為單一錯誤中斷,而是繼續執行其他 test cases。
執行完成後,可以查看 data/eval_runs/:
ls data/eval_runs
應該會看到類似:
eval_run_20260906_103000.json
也可以用 python3 -m json.tool 檢查輸出的 JSON:
python3 -m json.tool data/eval_runs/eval_run_20260906_103000.json
實際檔名要換成你當次產生的檔案名稱。
裡面會包含每一題的:
case_id
input
expected
grading_method
task_type
status
actual
trace_session_id
error
Day 11 的 runner 會對每一筆成功執行的 test case 呼叫:
save_trace(result.trace)
所以執行 batch run 後,Trace Viewer 裡應該會多出多筆 session。
啟動 Trace Viewer:
streamlit run trace_viewer_app.py
接著可以選擇某一筆 session,查看該 test case 的執行過程。
這表示我們已經把兩個系統接起來:
Batch Evaluation Runner
-> Agent Runner
-> Trace Storage
-> Trace Viewer
雖然現在還沒有評分,但已經可以批次產生可追蹤的 Agent 執行紀錄。
今天很容易想順手把 pass / fail 做完。
例如:
passed = expected in actual
但依照這個系列的節奏,Day 11 只處理 batch execution。
原因是「執行」和「評分」是兩個不同責任。
Batch Runner 負責:
把 test cases 跑完,收集 Agent output。
Evaluator 負責:
根據 grading_method 判斷 output 是否符合 expected。
如果今天把兩者混在一起,程式會很快變大,也不利於後面擴充不同 grading method。
所以 Day 12 會專門處理 evaluator,先實作最簡單的:
今天完成後,系統具備:
evals/cases.json。SimpleAgent。trace_session_id 回到 Trace Viewer 查看執行過程。目前還沒有:
Day 10 建立了 eval dataset。
Day 11 則讓 dataset 真的跑起來。
今天最重要的流程是:
cases.json
-> evals.runner
-> SimpleAgent.run()
-> AgentResult
-> save_trace()
-> eval_run_*.json
這表示我們已經從「手動測試」前進到「批次執行」。
雖然目前還沒有判斷對錯,但這一步很重要。
因為只有先穩定收集每題的 Agent output,後面才有辦法做自動評分、成功率統計與錯誤分析。
Day 12 會實作最簡單的自動評分。
下一篇會新增 evaluator,先處理兩種 grading method:
exact_match
contains
到時候 evaluation result 就會從:
{
"status": "completed",
"actual": "The result is 3780"
}
進一步變成:
{
"status": "completed",
"actual": "The result is 3780",
"passed": true,
"failure_reason": null
}
這樣平台就會開始回答第二週最關心的問題:
Agent 到底有沒有完成任務?